Phase 3: Machine Learning Lesson 4 of 6

Unsupervised Learning:
Clustering

What if nobody told you the answers? Classification and regression both need labels. But some of the most interesting machine learning problems start with no labels at all. Clustering lets a model discover hidden structure in your data entirely on its own.

You will learn
What unsupervised learning is and why it matters
How K-Means clustering works step by step
How to choose the right number of clusters
Real-world applications of clustering
Running K-Means in Scikit-learn

Learning without labels

Supervised learning, which you covered in the last two lessons, requires labelled examples. Someone has to go through the data and write down the correct answer for each row. That is expensive, slow, and sometimes impossible. You cannot label every customer, every web page, every image on the internet.

Supervised learning

Has labels. The model learns from input-output pairs. You know the correct answer for every training example.

email → spam or not spam
image → cat or dog
house features → price
Unsupervised learning

No labels. The model finds patterns in raw data without being told what to look for. The structure emerges from the data itself.

customers → natural groups
articles → topics
genes → expression patterns

Unsupervised learning is what you turn to when you have data but no answers. When a streaming service wants to understand its user base without manually categorising millions of accounts. When a biologist wants to see which genes behave similarly without having any prior theory about the groupings. When a retailer wants to understand natural buying patterns before designing targeted promotions.

Clustering is the most widely used unsupervised technique. The goal is simple: group data points so that points in the same cluster are similar to each other and different from points in other clusters.

Analogy

You have just inherited a huge box of old photographs from a relative. No names, no dates, no labels. But as you spread them across a table, you start to notice things. These 30 photos all seem to feature the same family at a beach house. These 20 are all winter holidays. These 15 all look like work events. You are clustering: grouping similar things together without anyone telling you how many groups to make or what to call them. K-Means does the same thing with numbers.

K-Means: the simplest clustering algorithm

K-Means is the standard starting point for clustering. It is fast, interpretable, and works surprisingly well on a huge range of problems. The name tells you two things: K is the number of clusters you specify in advance, and "Means" refers to the centroid (average position) of each cluster.

K-Means clustering: three iterations toward convergence
Step 1: Initialise
Place K centroids randomly
Step 2: Assign
Each point gets nearest centroid
Step 3: Update
Move centroid to cluster mean

Steps 2 and 3 repeat until the centroids stop moving. Convergence usually happens in 10 to 20 iterations. The final centroids define the cluster centres. Every new data point is assigned to whichever centroid it is closest to.

The algorithm is deceptively simple, but it makes one important assumption: clusters are roughly spherical and similarly sized. If your data has elongated, irregular, or very differently sized groups, K-Means will struggle. For those cases, alternatives like DBSCAN (which finds clusters of arbitrary shape) or Gaussian Mixture Models (which assign probabilities rather than hard labels) work better.

Choosing the right K

The biggest practical challenge with K-Means is that you have to specify K before you run the algorithm. How many clusters are there in your data? If you already knew that, you would not need clustering. You are trying to discover structure, not impose it.

The elbow method is the standard way to make this choice. You run K-Means with K ranging from 1 to 10 (or more), and for each K you record the inertia: the total squared distance from every point to its assigned centroid. As K increases, inertia always decreases. But the rate of decrease slows dramatically after the "true" number of clusters. You look for the elbow in the curve.

Elbow method: inertia vs number of clusters
Number of clusters (K) Inertia 1 2 3 4 5 6 7 8 9 Elbow = K=3

The curve drops steeply from K=1 to K=3, then starts to flatten. The elbow at K=3 suggests three is the natural number of clusters in this dataset. Adding more clusters beyond this point gives diminishing returns in how well the model fits the data.

The elbow is often a judgment call rather than a crisp mathematical answer. If the curve is smooth with no clear bend, the silhouette score is another metric worth checking. It measures how tightly each point fits within its own cluster compared to its distance from neighbouring clusters. Values close to 1 mean good clusters; values near 0 mean overlapping clusters.

Where clustering shows up in the real world

👥
Customer segmentation
Retail and e-commerce companies cluster customers by purchase history, browsing patterns, and demographics to create targeted campaigns and personalised recommendations.
📰
Document clustering
News aggregators and search engines group articles about the same topic together without anyone labelling them. Topic modelling is a related technique that also reveals themes.
🧬
Genomics and biology
Researchers cluster genes with similar expression profiles across conditions to identify gene families or discover which genes are regulated together.
🔍
Anomaly detection
Data points that do not belong to any cluster (or are very far from any centroid) are outliers. This catches fraud, network intrusions, or manufacturing defects.

K-Means in Scikit-learn

The Scikit-learn API for clustering is the same pattern you have seen for everything else. Create the object, fit it, get the labels out. The main difference is that fit does not need a y (target) argument, because there are no labels.

Python customer_clustering.py
import numpy as np
import matplotlib.pyplot as plt
from sklearn.cluster import KMeans
from sklearn.preprocessing import StandardScaler

# Simulate customer data: annual spend and visit frequency
np.random.seed(42)
spend = np.random.normal(loc=[500, 2000, 8000], scale=[100, 300, 800], size=(100, 3)).flatten()
visits = np.random.normal(loc=[2, 8, 20], scale=[1, 2, 4], size=(100, 3)).flatten()
X = np.column_stack([spend, visits])

# Scale the features (important for distance-based algorithms)
scaler = StandardScaler()
X_scaled = scaler.fit_transform(X)

# Fit K-Means with 3 clusters
kmeans = KMeans(n_clusters=3, random_state=42, n_init=10)
labels = kmeans.fit_predict(X_scaled)

# Summarise each cluster
for k in range(3):
    mask = labels == k
    print(f"Cluster {k}: {mask.sum()} customers | "
          f"avg spend=${X[mask,0].mean():.0f} | "
          f"avg visits/yr={X[mask,1].mean():.1f}")
Output
Cluster 0: 100 customers | avg spend=$499 | avg visits/yr=2.1
Cluster 1: 100 customers | avg spend=$7,984 | avg visits/yr=20.2
Cluster 2: 100 customers | avg spend=$1,997 | avg visits/yr=8.0

Three distinct customer segments emerge automatically: low-spend infrequent visitors, mid-range regular customers, and high-value loyal shoppers. No human had to categorise a single customer. The model found the natural groupings in the data.

Always scale before clustering

K-Means uses Euclidean distance. If one feature is measured in thousands (annual spend) and another in single digits (monthly visits), the distance will be dominated by the large-scale feature. Always run StandardScaler first so each feature contributes equally to the distance calculation. Skipping this step is one of the most common mistakes beginners make with clustering.

Hands-on activity

Cluster your own dataset and name the groups

You will run K-Means on a real dataset of mall customers (spend vs income), use the elbow method to choose K, visualise the results, and then do the most human part of the job: naming the clusters based on what you see.

01 Open the Lesson 3.4 Colab notebook. Load the Mall Customers dataset and plot Annual Income vs Spending Score as a scatter chart.
02 Run the elbow method for K values 1 through 10. Where is the elbow? What does it suggest about the natural structure of this data?
03 Fit KMeans with your chosen K and plot the clusters in different colours, with centroids marked as stars. Does the visual make intuitive sense?
04 For each cluster, compute the mean income and mean spending score. Give each cluster a descriptive name (e.g., "Careful savers," "High earners, low spenders," "Impulsive shoppers").
05 Try running the same code WITHOUT StandardScaler first. Does the clustering change? Why or why not?
Your Notes
Studying independently? Write your thoughts or answers below. Notes save automatically to your browser.
Practice Notebook
Run this lesson's code live in Google Colab
All examples + challenge exercises · Free GPU included · No setup required
Open In Colab
Progress
Done with this lesson?
Mark it complete to track your progress.